Goto

Collaborating Authors

 kohler and langer


Covariate-dependent Graphical Model Estimation via Neural Networks with Statistical Guarantees

arXiv.org Machine Learning

Graphical models are widely used in diverse application domains to model the conditional dependencies amongst a collection of random variables. In this paper, we consider settings where the graph structure is covariate-dependent, and investigate a deep neural network-based approach to estimate it. The method allows for flexible functional dependency on the covariate, and fits the data reasonably well in the absence of a Gaussianity assumption. Theoretical results with PAC guarantees are established for the method, under assumptions commonly used in an Empirical Risk Minimization framework. The performance of the proposed method is evaluated on several synthetic data settings and benchmarked against existing approaches. The method is further illustrated on real datasets involving data from neuroscience and finance, respectively, and produces interpretable results.


Predicting path-dependent processes by deep learning

arXiv.org Machine Learning

In this paper, we investigate a deep learning method for predicting path-dependent processes based on discretely observed historical information. This method is implemented by considering the prediction as a nonparametric regression and obtaining the regression function through simulated samples and deep neural networks. When applying this method to fractional Brownian motion and the solutions of some stochastic differential equations driven by it, we theoretically proved that the $L_2$ errors converge to 0, and we further discussed the scope of the method. With the frequency of discrete observations tending to infinity, the predictions based on discrete observations converge to the predictions based on continuous observations, which implies that we can make approximations by the method. We apply the method to the fractional Brownian motion and the fractional Ornstein-Uhlenbeck process as examples. Comparing the results with the theoretical optimal predictions and taking the mean square error as a measure, the numerical simulations demonstrate that the method can generate accurate results. We also analyze the impact of factors such as prediction period, Hurst index, etc. on the accuracy.


Posterior and variational inference for deep neural networks with heavy-tailed weights

arXiv.org Machine Learning

We consider deep neural networks in a Bayesian framework with a prior distribution sampling the network weights at random. Following a recent idea of Agapiou and Castillo (2023), who show that heavy-tailed prior distributions achieve automatic adaptation to smoothness, we introduce a simple Bayesian deep learning prior based on heavy-tailed weights and ReLU activation. We show that the corresponding posterior distribution achieves near-optimal minimax contraction rates, simultaneously adaptive to both intrinsic dimension and smoothness of the underlying function, in a variety of contexts including nonparametric regression, geometric data and Besov spaces. While most works so far need a form of model selection built-in within the prior distribution, a key aspect of our approach is that it does not require to sample hyperparameters to learn the architecture of the network. We also provide variational Bayes counterparts of the results, that show that mean-field variational approximations still benefit from near-optimal theoretical support.


Posterior concentrations of fully-connected Bayesian neural networks with general priors on the weights

arXiv.org Machine Learning

Bayesian approaches for training deep neural networks (BNNs) have received significant interest and have been effectively utilized in a wide range of applications. There have been several studies on the properties of posterior concentrations of BNNs. However, most of these studies only demonstrate results in BNN models with sparse or heavy-tailed priors. Surprisingly, no theoretical results currently exist for BNNs using Gaussian priors, which are the most commonly used one. The lack of theory arises from the absence of approximation results of Deep Neural Networks (DNNs) that are non-sparse and have bounded parameters. In this paper, we present a new approximation theory for non-sparse DNNs with bounded parameters. Additionally, based on the approximation theory, we show that BNNs with non-sparse general priors can achieve near-minimax optimal posterior concentration rates to the true model.


Estimation of a regression function on a manifold by fully connected deep neural networks

arXiv.org Machine Learning

Deep neural networks (DNNs) are built of multiple layers and learn sequentially multiple levels of representation and abstraction by performing a nonlinear transformation on the data. The approach has proven itself to work incredibly well in practice, like for speech (Graves et al. (2013)) and image recognition (Krizhevsky et al. (2017)), or game intelligence (Silver et al. (2016)). But, unfortunately, the procedure is not well understood. Recently, several researchers tried to explain the performance of DNNs from a theoretical point of view. Results concerning the approximation power of DNNs were shown in Montufar (2014), Eldan and Shamir (2016), Yarotsky (2017), Yarotsky and Zhevnerchuck (2020), Langer (2021b) and Lu et al. (2020). Beside this, quite a few articles try to answer the question about why neural networks perform well on unknown new data sets (cf., e.g., Bauer and Kohler (2019), Schmidt-Hieber (2020), Kohler and Langer (2020), Kohler, Krzyżak, and Langer (2019), Langer (2021a), Imaizumi and Fukumizu